Papers with Offline preference optimization methods
Intrinsic Mutual Information as a Modulator for Preference Optimization (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for offline preference optimization involve additional hyperparameter tuning, resulting in substantial time overhead. |
| Approach: | They propose a lightweight framework for offline preference optimization that leverages hyperparameter modulation to decouple preference contributions. |
| Outcome: | The proposed framework achieves superior performance over existing methods while reducing training overhead by more than 15%. |
Adaptive Preference Optimization with Uncertainty-aware Utility Anchor (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Offline preference optimization methods are efficient for large language models (LLMs) alignment. |
| Approach: | They propose an offline preference optimization framework that estimates uncertainties from preference data . the method enables training even in scenarios where the data is unpaired . |
| Outcome: | The proposed method enables training even in scenarios where the data is unpaired . |